Papers with Red Teaming
Red Teaming Visual Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | VLMs (Vision-Language Models) can be induced to generate harmful or inaccurate content through specific test cases. |
| Approach: | They propose a red teaming dataset which encompasses 12 subtasks under 4 primary aspects (faithfulness, privacy, safety, fairness) this dataset is the first to benchmark current VLMs in terms of these 4 aspects . |
| Outcome: | The proposed dataset shows that 10 open-source VLMs struggle with red teaming in different degrees and have up to 31% performance gap with GPT-4V. |